Papers with acoustic model
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation (N19-1)
Copied to clipboard
| Challenge: | Pre-training on high-resource automatic speech recognition (ASR) tasks improves ST performance even when source language is low-resourced. |
| Approach: | They propose a method to improve direct speech-to-text translation when source language is low-resource . they pre-train model on high-res automatic speech recognition task and fine-tune parameters for ST . |
| Outcome: | The proposed approach improves Spanish English ST even when the source language is low-resource . the pre-trained encoder accounts for most of the improvement, the authors show . |
Scaling Under-Resourced TTS: A Data-Optimized Framework with Advanced Acoustic Modeling for Thai (2025.acl-industry)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) systems are limited by limited data and linguistic complexities. |
| Approach: | They propose a data-optimized framework with an advanced acoustic model to build high-quality TTS systems for low-resource scenarios. |
| Outcome: | The proposed framework enables zero-shot voice cloning and improved performance across diverse client applications, including finance, healthcare, education, and law. |
An Application for Building a Polish Telephone Speech Corpus (L18-1)
Copied to clipboard
| Challenge: | Specifically, we describe a tool designed to improve our Automatic Speech Recognition system performance. |
| Approach: | They propose to build a tool for speech corpus collection of a specific domain content. |
| Outcome: | The proposed tool can be used to gather 63 hours of speech recordings across several domains and achieve lower WER in two grammar-based speech recognition tasks. |
Automatic Speech Recognition for Gascon and Languedocian Variants of Occitan (2024.lrec-main)
Copied to clipboard
Iñigo Morcillo, Igor Leturia, Ander Corral, Xabier Sarasola, Michaël Barret, Aure Séguier, Benaset Dazéas
| Challenge: | a new system for automatic speech recognition is being developed for two main Occitan dialects . the difficulty lies in the fact that Occitian is a less-resourced language . |
| Approach: | They propose to develop an automatic speech recognition system for two Occitan dialects . they use Kaldi, acoustic models, and Whisper to create a model from corpora . |
| Outcome: | The proposed system is based on Kaldi and Whisper for two main Occitan dialects . the system is more robust than previous systems, and the results are promising . |
Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local Modeling (2023.emnlp-main)
Copied to clipboard
| Challenge: | Singing Voice Synthesis (SVS) synthesizes pleasing vocals based on music scores and lyrics . current acoustic models ignore the significance of local modeling within the sequence and the hard-to-synthesize parts in the predicted mel-spectrogram . |
| Approach: | They propose a method to enhance local modeling in the acoustic model by focusing on phoneme tokens located before and after the phoneme. |
| Outcome: | The proposed method improves local modeling in the acoustic model by focusing on the hard-to-synthesize parts of the predicted mel-spectrogram. |
Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to learn the transfer from speech to text are unexplored . how to solve the representation discrepancy of speech and text is unexplorable . |
| Approach: | They propose a cooperative acoustic and linguistic representation learning method to fuse and utilize contextual information of speech and text. |
| Outcome: | The proposed method outperforms existing methods on low-resource speech recognition. |
PronouncUR: An Urdu Pronunciation Lexicon Generator (L18-1)
Copied to clipboard
| Challenge: | acoustic modeling, large text data and a pronunciation lexicon are the bottlenecks for speech recognition systems for resource scarce languages. |
| Approach: | They propose a grapheme-to-phoneme conversion tool that generates a pronunciation lexicon from a list of Urdu words. |
| Outcome: | The proposed tool predicts pronunciation of words using a LSTM-based model trained on a handcrafted expert lexicon of around 39,000 words and shows an accuracy of 64% upon internal evaluation. |
Automatic Speech Recognition in Sanskrit: A New Speech Corpus and Modelling Insights (2021.findings-acl)
Copied to clipboard
| Challenge: | In this paper, we propose the first large scale study of automatic speech recognition in Sanskrit . we focus on the impact of unit selection in San's ASR systems . |
| Approach: | They propose a large scale study of automatic speech recognition in Sanskrit . they propose syllable level unit selection that captures character sequences . |
| Outcome: | The proposed model captures character sequences from one vowel in the word to the next vowela. |
Developing Resources for Automated Speech Processing of Quebec French (2020.lrec-1)
Copied to clipboard
| Challenge: | acoustic models for automatic segmentation of Quebec French are not available for all languages . linguistic resources are developed to perform phonetic annotations in Quebec French . physical characteristics of speech can be observed in the production of sounds . |
| Approach: | They propose to use a French lexicon to train automatic QF segmentation models . they adapt existing pronunciation dictionary and acoustic model from existing ones . |
| Outcome: | The proposed tools perform the full process of speech segmentation in Quebec French. |
ASR for Documenting Acutely Under-Resourced Indigenous Languages (L18-1)
Copied to clipboard
| Challenge: | Automatic speech recognition (ASR) has not been widely explored as a tool for documenting endangered languages. |
| Approach: | They propose to use automatic speech recognition (ASR) to bootstrap new data to improve the acoustic model. |
| Outcome: | The proposed system improves the model for a polysynthetic language with few audio and text resources. |
A Romanization System and WebMAUS Aligner for Arabic Varieties (2022.lrec-1)
Copied to clipboard
Jalal Al-Tamimi, Florian Schiel, Ghada Khattab, Navdeep Sokhey, Djegdjiga Amazouz, Abdulrahman Dallak, Hajar Moussa
| Challenge: | The WebMAUS 1 is a suite of webservices that is free for academic users that processes 42 languages and language varieties. |
| Approach: | They propose to develop an Arabic variety-independent romanization system that aims to homogenize and simplify the romanization of the Arabic script. |
| Outcome: | The proposed system is based on the existing Arabic variety-independent WebMAUS services. |
Automatic Speech Recognition for Uyghur through Multilingual Acoustic Modeling (2020.lrec-1)
Copied to clipboard
| Challenge: | Low-resource languages suffer from lower performance of Automatic Speech Recognition (ASR) due to the lack of data. |
| Approach: | They propose to use Turkish as donor language to train acoustic models using multilingual training to achieve more context coverage. |
| Outcome: | The proposed system performs better with multilingual training for the under-resourced Uyghur language. |
On Construction of the ASR-oriented Indian English Pronunciation Dictionary (2020.lrec-1)
Copied to clipboard
| Challenge: | Indian English (IE) has distinctive characteristics, especially phonologically, from other varieties of English. |
| Approach: | They build a small IE spontaneous speech corpus and use a linguistically-guided IE pronunciation dictionary to apply it to IE. |
| Outcome: | The proposed system performs better on IE spontaneous speech data than the one trained with CMUdict. |
Multimodal In-context Learning for ASR of Low-resource Languages (2026.findings-acl)
Copied to clipboard
| Challenge: | In-context learning with large language models addresses this limitation, but prior work focuses on high-resource languages covered during training and text-only settings. |
| Approach: | They propose to use multimodal ICL to learn unseen languages with multimodal learning to improve ASR in large language models. |
| Outcome: | The proposed model outperforms existing models on unseen languages with multimodal ICL (MICL) and cross-lingual transfer learning matches or outperformed models without using target-language data. |